Characters, Glyphs and Beyond
نویسندگان
چکیده
The distinction between characters and glyphs is a fundamental issue of computing. This talk aims in giving a new definition of these notions. We first review and comment the definitions given in various standards. Then we give and explain our own definitions. We consider that the Unicode character model is lacunary and formulate a proposal for adding supplementary information and obtaining thus “rich Unicode characters.” We illustrate our arguments with many examples, taken from various writing systems. The distinction between characters and glyphs is currently a very popular issue. The complexity of this issue is, in some sense, related to the fact that computer systems have been build by engineers not very proficient in linguistics, and interested only in the English language. Exploring non-latin writing systems one realizes what has not been clear from the beginning: that modelizing written language is not a trivial task, and that it is fundamental to all exchange and processing of textual information. Let us start the exploration of this universe by giving some definitions of the terms we are using. Let us see how the terms “character” and “glyph” are defined. According to ISO 9541 [6] released in 1991, a “glyph” is “a recognizable abstract graphic symbol which is independent of any specific design,” while a “glyph image” is “an image of a glyph, as obtained from a glyph representation diplayed on a presentation surface,” where “glyph representation” is “the glyph shape and glyph metrics associated with a specific glyph in a font resource.” We may argue if this distinction between “abstract glyph” and “concrete glyph” is necessary, but this is how ISO 9541 defines these. According to W3C (quoting “A Character Model for the World Wide Web” by Martin Drst and others [2]), a character is “the smallest component of written language that has semantic values; refers to the abstract meaning and/or shape.” We find this definition quite vague since everything we perceive may or may not have semantic value, depending on our culture, context and even mood. . .We all know that Unicode is full of inconsistencies, because of its requirement to be compatible with legacy encodings. Has this definition been made to encompass Unicode
منابع مشابه
Providing some UTF-8 support via inputenc
3 Mapping characters — based on font (glyph) encodings 11 3.1 About the table itself . . . . . . . . . . . . . . . . . . . . . . . . . 12 3.2 The mapping table . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 3.3 Notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 3.4 Mappings for OT1 glyphs . . . . . . . . . . . . . . . . . . . . . . . 24 3.5 Mappings for OMS g...
متن کاملLearning Chinese Word Representations From Glyphs Of Characters
In this paper, we propose new methods to learn Chinese word representations. Chinese characters are composed of graphical components, which carry rich semantics. It is common for a Chinese learner to comprehend the meaning of a word from these graphical components. As a result, we propose models that enhance word representations by character glyphs. The character glyph features are directly lea...
متن کاملGeneration of Glyphs for Conveying Complex Information, with Application to Protein Representations
We present a method to generate glyphs which convey complex information in graphical form. A glyph has a linear geometry which is specified using geometric operations, each represented by characters nested in a string. This format allows several glyph strings to be concatenated, resulting in more complex geometries. We explore automatic generation of a large number of glyphs using a genetic alg...
متن کاملOmega Becomes a Texteme Processor
The distinction between “characters” and “glyphs” is a rather new issue in computing, although the problem is as old as humanity: our species turns out to be a writing one because, amongst other things, our brain is able to interpret images as symbols belonging to a given writing system. Computers deal with text in a more abstract way. When we agree that, in computing, all possible “capital A” ...
متن کاملΩTimes and ΩHelvetica Fonts Under Development: Step One
ΩTimes and ΩHelvetica will be public domain virtual Timesand Helvetica-like fonts based upon real PostScript fonts, which we call “Glyph Containers”. They will contain all necessary characters for typesetting efficiently (that is, with TEX quality) in all languages and systems using the Latin, Greek, Cyrillic, Arabic, Hebrew and Tifinagh alphabets and their derivatives. All Unicode characters w...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2004